Launching the Linux release and noticed in the logs:
Directories:User Directory: /home/bisby/Grayjay
And there is a directory there now. I absolutely hate having stuff automatically create anything in my home directory like this. Ideally, this should be following XDG directory guidelines on linux: https://specifications.freedesktop.org/basedir-spec/latest/
Complete aside here: I used to do work with amputees and prosthetics. There is a standardized test (and I just cannot remember the name) that fits in a briefcase. It's used for measuring the level of damage to the upper limbs and for prosthetic grading.
Basically, it's got the dumbest and simplest things in it. Stuff like a lock and key, a glass of water and jug, common units of currency, a zipper, etc. It tests if you can do any of those common human tasks. Like pouring a glass of water, picking up coins from a flat surface (I chew off my nails so even an able person like me fails that), zip up a jacket, lock your own door, put on lipstick, etc.
We had hand prosthetics that could play Mozart at 5x speed on a baby grand, but could not pick up a silver dollar or zip a jacket even a little bit. To the patients, the hands were therefore about as useful as a metal hook (a common solution with amputees today, not just pirates!).
Again, a total aside here, but your comment just reminded me of that brown briefcase. Life, it turns out, is a lot more complex than we give it credit for. Even pouring the OJ can be, in rare cases, transcendent.
I wonder if the new drug of choice is actually technology. In some ways I think that the addiction to technology has some similar mellowing effects as drugs. Some research indicates that smartphone addiction is also related to low self-esteem and avoidant attachment [1] and that smartphones can become an object of attachment [2]. The replacement of drugs by technology is not surprising as it significantly strengthens technological development especially as it is already well past the point of diminishing returns for improving every day life.
1. https://www.sciencedirect.com/science/article/abs/pii/S07475...
2. https://www.sciencedirect.com/science/article/abs/pii/S07475...
Efficiency is now key.
~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task.
We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective experience) between 5 second and 5 minutes to solve the task. (So i'd argue a human is at 0.03USD - 1.67USD per puzzle at 20USD/hr, and they include in their document an average mechancal turker at $2 USD task in their document)
Going the other direction: I am interpreting this result as human level reasoning now costs (approximately) 41k/hr to 2.5M/hr with current compute.
Super exciting that OpenAI pushed the compute out this far so we could see he O-series scaling continue and intersect humans on ARC, now we get to work towards making this economical!
Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far.
A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning.
It's obvious to everyone that these models can't perform as well as humans on everyday tasks despite blowout scores on the hardest tests we give to humans. Yet nobody could quantify exactly the ways the models were deficient. ARC is the best effort in that direction so far.
We don't need more "hard" benchmarks. What we need right now are "easy" benchmarks that these models nevertheless fail. I hope Francois has something good cooked up for ARC 2!
> I knew this was impossible, because…
There’s an easier tell. It’s impossible because you can’t to get Google to help you at all about any account issues, never mind them being as proactive as to call you.
In other words if Google call you, it’s not Google.
It’s slightly depressing that there are probably more fake Google support staff than real ones.
This article is fascinating. But what's on display here is less of a nefarious plan from Spotify to replace famous Katy Perry with AI - instead we get to see something much more specific: a behind-the-scenes of how those endless chill/lo-fi/ambient playlists get created.
Which is something I've always wondered! How does the Lofi Girl channel on Youtube always have so much new music from artists I have never heard from?
The answer is surprising: real people and real instruments! (At least at the time of writing). Third-party stock music ("muzak") companies hiring underemployed jazz musicians to crank out a few dozen derivative songs every day to hack the algorithm.
> “Honestly, for most of this stuff, I just write out charts while lying on my back on the couch,” he explained. “And then once we have a critical mass, they organize a session and we play them. And it’s usually just like, one take, one take, one take, one take. You knock out like fifteen in an hour or two.” With the jazz musician’s particular group, the session typically includes a pianist, a bassist, and a drummer. An engineer from the studio will be there, and usually someone from the PFC partner company will come along, too—acting as a producer, giving light feedback, at times inching the musicians in a more playlist-friendly direction.”
I think there's an easy and obvious thing we can do - stop listening to playlists! Seek out named jazz artists. Listen to your local jazz station. Go to jazz shows.
Good times. I was the developer at Microsoft who designed the Xbox 360 hardware security, wrote all the boot loaders, and the hypervisor code.
Note to self: you should have added random delays before and after making the POST code visible on the external pins.
There exists a small percentage of men who will go absolutely savage on somebody for stealing from them, and the existence of those people is probably a bigger crime deterrent than the police.
So I say, shine on you crazy air tag tracking vigilante diamonds.
> We cannot find qualified applicants.
Your company is incompetent. I've applied to hundreds of companies like yours within Huntsville, AL in the past year, rejected or ghosted all the time.
Defense morons will talk about how hard their work is and how they can't find anyone to do it. Completely skip over how prevalent affirmative action is in their hiring process; who were you guys interviewing in 2020? Why is the defense small business base completely dominated by veterans who stack 10% disability ratings and minorities with a preferred SBA sticker on their website?
Complete joke of an industry.
So many contradictions!
You pay well, but not so much.
You search for qualified applicants but can hire a student.
You require linear algebra but ok with technical writer.
Looks like your managers don't know who they need to hire or don't want to really hire.
My employer cannot hire H-1Bs.
You must be a US citizen to work for my company. No "US Persons" (visa holders) or foreigners allowed.
You have to be eligible for a Secret security clearance. You don't have to get one if you don't want to as there is usually plenty of uncleared work to go around, but you have to be eligible in case that goes away and we need to put you in for a clearance.
We cannot find qualified applicants.
I've had this conversation many times on HN so here are some preemptive responses:
No, we don't make weapons for the military. Well, we do but not my part of the company. The most harmful thing the products I build do is quantify in precise detail how climate change is dooming us all.
No, our positions aren't ghost positions.
Yes, we are willing to train someone who is motivated. We won't re-teach linear algebra to a developer applicant but we will pay a tech writer to go to school nights/weekends to get a degree in engineering (me, I did that).
Yes, we have extensive high school and college work-study/internships and participants make $72k/yr. with full benefits for the duration of the program. That pipeline is actually successful.
No, you can't work remotely. You (even programmers!) have to touch the things we build in order to build them and nobody has an ISO certified clean room in their house.
Yes, we pay well.
No, we don't pay as much as Meta. We build components for satellites that have been sold to space agencies and purchased by various departments/ministries of the environment, not your personal information to advertisers-- one party has more money to spend than the other.
We have shortages in mech/EE/Aero, shortages in software, and critical shortages in engineering technicians.
One issue is that we expect programmers to remember linear algebra and have more than the ability to shovel frameworks on top of each other until a phone app comes out the other side.
I've had an MX Master mouse (the "2" for maybe 8-9 years then the "3" for 2-3 years now) and love it. Great performance, great battery life, fantastic design and feel. On Windows I definitely do not love the 150Mb program to manage it (surely sending a torrent of unnecessary telemetry data back to Logitech.
I found Solaar a couple months ago after getting repeatedly frustrated with bluetooth connection issues. It really is exactly what it needs to be. Better interface than Logitech's, simple, lightweight. Devs have my thanks; what a great show of the goodness of open source software.
A supposed shortage of qualified US applicants for tech jobs, especially software developers, doesn't jibe with the huge numbers of US developers currently looking for work, including highly experienced older workers suffering from age discrimination.
I'd be surprised if more than 5-10% of H-1B positions are ones where the hiring company has even looked for US applicants.
Ah, classic regulatory theater. The administration, after 4 years of not introducing these changes, is now suddenly scrambling to roll them out. They’re dropping them right before a major transition, with an implementation timeline conveniently set for after the transition.
It’s a clever little maneuver. When the inevitable reversal happens, they can show up at fundraising galas telling donors, “We tried! We were so close! It’s just those baddies who always come along and pull the rug.”
It seems very odd that the article seems to be measuring the information content of specific tasks that the brain is doing or specific objects that it is perceiving. But the brain is a general-purpose computer, not a speed-card computer, or English text computer, or binary digit computer, or Rubik's cube computer.
When you look at a Rubik's cube, you don't just pick out specific positions of colored squares relative to each other. You also pick up the fact that it's a Rubik's cube and not a bird or a series of binary digits or English text. If an orange cat lunged at your Rubik's cube while you were studying it, you wouldn't process it as "face 3 has 4 red squares on the first row, then an orange diagonal with sharp claws", you'd process it as "fast moving sharp clawed orange cat attacking cube". Which implies that every time you loom at the cube you also notice that it's still a cube and not any of the millions of other objects you can recognize, adding many more bits of information.
Similarly, when you're typing English text, you're not just encoding information from your brain into English text, you're also deciding that this is the most relevant activity to keep doing at the moment, instead of doing math or going out for a walk. Not to mention the precise mechanical control of your muscles to achieve the requisite movements, which we're having significant trouble programming into a robot.
> While we gained some traction with some startups, we ultimately weren't able to build a hyper-growth business.
>"We then decided to pivot into a vertical B2B SaaS AI product because we felt we could use the breakthroughs in Gen AI to solve previously unsolvable problems, but after going through user interviews and the sales cycle for many different ideas, we haven't been able to find enough traction to make us believe that we were on the right track to build a huge business."
How many wonderful niche products would be around if their owners had tried for a small business instead of a 'huge business' with 'hyper growth'?
This shouldn't be controversial. FTC isn't banning these fees, it is only requiring merchants to disclose the fees. Why would anyone be against that?
Another good rule is click-to-cancel. Just a couple of days ago I logged into my Dish Network account to cancel it (after they hiked prices). There is no way to cancel online. There is no way to cancel via chat. You have to call. As soon as you call you're told the wait time is over 45 minutes. There is no call back option. Why should a consumer have to be on the phone for 45 minutes to cancel? (Typically they will drop the call after 45 minutes and you have to call again.) If you call Dish to sign up service the wait time is 0 minutes: they answer immediately. If you then tell that you're actually calling to cancel, they forward you to the cancellation number with the wait. This is an abusive business practice, and banning it should not be controversial.
This is wonderful!!! Generalizing here but we really do take the moon for granted.
I bought a 'big ass telescope' a few years ago in an effort to bootstrap a hobby that I'd flirted with for decades but never really committed to. It's a Celestron 11" SCT and I really had no idea what I was getting into. When I think of space I think of things that are really small in the night sky, planets, galaxies, nebula...(turns out most of them aren't *that* small and I overshot the targets I had in mind)
I kept trying to photo galaxies and star clusters and all of these exotic things but had a bunch of trouble with tracking with long exposures. Out of frustration I ended up just pointing it at the boring ol' moon to at least get used to the equipment and workflows.
I fell in love with Luna.
The magnification of this scope really allowed me to explore the surface in a way I never had before. I got to know the 'map' and suddenly related to our celestial neighbor in a whole new way. It was also the very first image I was actually not embarrassed to share - https://imgur.com/a/t9b1Uug
I since then improved my knowledge and technical skill but the month of the moon at the end of 2021 was really pretty spectacular for me.
Passkeys are a terrible idea. They are security theater and a disaster for users waiting to happen.
Imagine you're on vacation and have lost your phone. You want to go to a cafe and log into a chat app, an email service, whatever to contact your family. In the world that passkey advocates want this is impossible via the passkey flow. If you can't authenticate via a primary device that contains your private key, you're f-ed. Service providers know this so of course they will provide recovery mechanisms. (Not consistent recovery mechanisms of course, each will have their own convoluted and likely to be broken ones). If the recovery mechanism allows for knowledge based recovery (challenge questions) then you're basically telling people that they need N passwords rather than one password. Maybe that challenge question is just a long password (recovery key). Maybe it requires access to another system, which you likely don't have in this circumstance. So you're either back to having passwords, or you're f-ed. Security theater or a disaster.
A service only has value if I can access it. I should be able to sit down at any computer in the world with the knowledge in my head and get access to any online service to which I subscribe.
Please visit people when they are still alive.
My aunt was alone for the last three years of her live, up to 94 years old. She had almost no visitors, and was not able to go out by herself. I went there about every week, and was always the only visitor.
Then came the funeral, with well over 400 people there, and around 200 people at the coffee table after the funeral. And I was like: Where the hell were you guys the past three years? She was alone (after her husband died) and most of you never visited her since (and no they were not living far away or unable to visit).
Go to the funeral yes, but don't wait until.
They stopped hiring because after raising over a billion dollars their value dropped 80%, they can't raise any more money, and after operating for 19 years they finally had a profitable quarter in Q2 2024. They are in "burn the furniture to heat the cabin" mode.
Is legalese not just the result of trying to use English as a programming language? Any time I try to write English (or other natlang) precisely and unambiguously and robust against adversarial interpretations, it comes out looking like legalese.
The last fairly technical career to get surprisingly and fully automated in the way this post displays concern about - trading.
I spent a lot of time with traders in early '00's and then '10's when the automation was going full tilt.
Common feedback I heard from these highly paid, highly technical, highly professional traders in a niche indusry running the world in its way was:
- How complex the job was - How high a quality bar there was to do it - How current algos never could do it and neither could future ones - How there'd always be edge for humans
Today, the exchange floors are closed, SWEs run trading firms, traders if they are around steer algos, work in specific markets such as bonds, and now bonds are getting automated. LLMs can pass CFA III, the great non-MBA job moat. The trader job isn't gone, but it has capital-C Changed and it happened quickly.
And lastly - LLMs don't have to be "great," they just have to be "good enough."
See if you can match the above confidence from pre-automation traders with the comments displayed in this thread. You should plan for it aggressively, I certainly do.
Edit - Advice: the job will change, the job might change in that you steer LLMs, so become the best at LLM steering. Trading still goes on, and the huge, crushing firms in the space all automated early and at various points in the settlement chain.
Is there some generalized law (yet) about unintended consequences? For example:
Increase fuel economy -> Introduce fuel economy standards -> Economic cars practically phased out in favour of guzzling "trucks" that are exempt from fuel economy standards -> Worse fuel economy.
or
Protect the children -> Criminalize activites that might in any way cause an increase in risk to children -> Best to just keep them indoors playing with electronic gadgets -> Increased rates of obesity/depression etc -> Children worse off.
As the article itself says: Hold big tech accountable -> Introduce rules so hard to comply with that only big tech will be able to comply -> Big tech goes on, but indie tech forced offline.
Nothing because I’m a senior and LLM’s never provide code that pass my sniff test, and it remains a waste of time.
I have a job at a place I love and get more people in my direct network and extended contacting me about work than ever before in my 20 year career.
And finally I keep myself sharp by always making sure I challenge myself creatively. I’m not afraid to delve into areas to understand them that might look “solved” to others. For example I have a CPU-only custom 2D pixel blitter engine I wrote to make 2D games in styles practically impossible with modern GPU-based texture rendering engines, and I recently did 3D in it from scratch as well.
All the while re-evaluating all my assumptions and that of others.
If there’s ever a day where there’s an AI that can do these things, then I’ll gladly retire. But I think that’s generations away at best.
Honestly this fear that there will soon be no need for human programmers stems from people who either themselves don’t understand how LLM’s work, or from people who do that have a business interest convincing others that it’s more than it is as a technology. I say that with confidence.
I don't worry about it, because:
1) I believe we need true AGI to replace developers.
2) I don't believe LLMs are currently AGI or that if we just feed them more compute during training that they'll magically become AGI.
3) Even if we did invent AGI soon and replace developers, I wouldn't even really care, because the invention of AGI would be such an insanely impactful, world changing, event that who knows what the world would even look like afterwards. It would be massively changed. Having a development job is the absolute least of my worries in that scenario, it pales in comparison to the transformation the entire world would go through.
> markets enforce efficiency, so it's not possible that a company can have some major inefficiency and survive
This just seems totally false on its face. If you've worked at the big guys you know they aren't magically smarter, they do very inefficient things frequently.
It's so intuitively false that I'd have to wonder about someone who thinks it's true.
This post is so interesting to me, esp. the build-vs-buy spectrum.
As Dan notes, a lot of software is just...not very good. It either isn't upfront with flaws (as in the case of the Postgres -> Snowflake tool), has too much scope, or is abstracted poorly. Finding things to buy/use (as in the case of open source) can often eat a lot more time than you anticipate.
I've been dipping my toes into the JS ecosystem, and I keep bumping into the fact that using mentally cheap signals of quality (such as stars or DL counts) almost never indicates the quality of the thing itself. Winners seem to be randomly chosen, almost! The only way to assess is to read the code and try integrating it in.
I'd go farther to argue that the larger an ecosystem/market is, the more untrustworthy it behaves as a whole, simply due to the size, and the types of people attracted to it who want to get influence/money. See also: appliances that everyone needs.
This is emblematic of the LLM race in general. We’re actively pressured to use co-pilot at work, and it’s crammed into every Microsoft product. I’m thankful that my iPhone is old enough not to use LLMs. Companies are afraid of being left behind in the new arms race, but that doesn’t actually mean that the technology actually present use-cases which most people need. (Worse are the meeting summaries or emails which are written by LLMs. The summaries are just not very good, and any sort of LLM writing is a tacit acknowledgement that people don’t really care what they are writing, an that no one is really reading that writing very carefully.)
This is called prototyping, which is a valuable part of the design process; some people call it "pathfinding".
These are all inputs into the design. But a design is still needed, of the appropriate size, otherwise you're just making things up as you go. You need to define the problem your are solving and what that solution is. Sometimes that's a 1-page doc without a formal review, sometimes it's many more pages with weeks of reviews and iterations with feedback.
Don't forget: "weeks of coding can save hours of planning" ;)
> challenged a group of Year 8 pupils to give up their smartphones completely for 21 days.
It was not a ban during school. It was complete phone abstinence. The result was that the kids got an entire additional hour of sleep! Perhaps this could be replicated just by putting phones away at night.